Papers with Polish language models
Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse (2025.acl-long)
Copied to clipboard
| Challenge: | specialized Polish language models are more effective at detecting harmful content than traditional methods. |
| Approach: | They propose a Polish-language dataset for erotic content detection that captures ambiguity, violence, and socially unacceptable behaviors. |
| Outcome: | The proposed dataset shows that specialized Polish language models achieve superior performance compared to multilingual alternatives, with transformer-based architectures showing particular strength in handling imbalanced categories. |
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing multilingual evaluation benchmarks focus on IR in the Polish language, but the Polish is a relatively new field due to the limited availability of Polish datasets. |
| Approach: | They propose to establish large-scale resources for IR in the Polish language and translate them into a new benchmark which includes 13 datasets. |
| Outcome: | The proposed benchmarks are based on 13 open IR datasets in Polish and are a pioneering development in this area. |